Back

Molecular Ecology Resources

Wiley

Preprints posted in the last 90 days, ranked by how well they match Molecular Ecology Resources's content profile, based on 171 papers previously published here. The average preprint has a 0.11% match score for this journal, so anything above that is already an above-average fit.

1
Nanopore sequencing of nested nrDNA barcodes reliably identifies orchid bees (Euglossini, Apidae)

Kolter, A.; Alvarado, M.; Roubik, D. W.; Eltz, T.

2026-08-26 zoology 10.64898/2026.08.25.747120 medRxiv
Top 0.1%
61.5%
Show abstract

Orchid bees (Euglossini, Apidae) are Neotropical insects whose species-level identification can depend on minute morphological characters, some difficult to see or analyze. In such cases, DNA barcoding may facilitate identification by comparing standardized DNA sequences with reference libraries. The mitochondrial cytochrome c oxidase I (COI) marker widely used in animals does not, however, provide uniform species-level resolution across bee lineages. We developed an adaptive-length nuclear ribosomal DNA (nrDNA) barcoding framework based on overlapping Nanopore-sequenced markers spanning approximately 500 to 5500 bp for 114 Euglossini species. By matching barcode length to specimen quality, material with varied preservation histories was processed within a single analysis. Leave-k-out validation with IDTAXA achieved more than 96% identification success for two longer barcodes, while performance was lower for the shortest. Combining barcode lengths within one reference library maintained high identification success, and confidence filtering reduced overclassification when species were absent from the reference library. For orchid bees, this framework permits affordable high-throughput identification and supports targeted taxonomic verification and revision. Combining adaptive barcode lengths in one analytical framework offers a general design principle for long-read reference-library construction. Its performance must now be tested in other groups.

2
A Draft Male Genome Assembly of the Slipper Lobster (Thenus australiensis) Reveals an XY System and a Validated Diagnostic Marker for Monosex Aquaculture.

Tran Nguyen, A. H.; Ha, G.-H.; Tran, D.-P.; Le, N. T.; Glendining, S.; Fitzgibbon, Q.; Herzig, V.; Luu, P.-L.; Ventura, T.

2026-06-29 genomics 10.64898/2026.06.24.734161 medRxiv
Top 0.1%
52.7%
Show abstract

The slipper lobster (Thenus australiensis) is rapidly emerging as a high-potential species for commercial aquaculture. Because females exhibit superior growth characteristics due to less frequent moulting after sexual maturity, developing monosex breeding strategies is highly desirable for industry profitability. However, the lack of genomic resources and early sex-identification tools has hindered this development. Here, we report the first draft male genome assembly for T. australiensis, generated using a combination of whole-genome shotgun sequencing, DArT-seq, and multi-tissue transcriptomics. The curated assembly spans 0.913 Gbp with high functional completeness (93.0% BUSCO), providing a robust repertoire of 30,100 protein-coding genes. Through k-mer subtraction and population-level DArT-seq genotyping, we provide definitive evidence that T. australiensis utilizes an XX/XY sex-determination system. Crucially, by identifying male-specific structural variations within a neo-Y locus, we developed a diagnostic PCR assay targeting a male-exclusive sequence. This 171 bp marker achieved 100% accuracy in phenotypic sex identification across wild-caught populations. Ultimately, these foundational genomic resources, combined with a highly reliable molecular sexing tool, provide the critical framework necessary for early sex sorting, broodstock management, and the commercial advancement of monosex slipper lobster farming.

3
Extended genomic regions flanking ultraconserved elements allow efficient species identification and intraspecific diversity assessment in coral

Mateos, A.; Cowman, P.; Bridge, T.; Yeoh, Y. K.; Bourne, D.; Sato, Y.

2026-08-28 genomics 10.64898/2026.08.25.743606 medRxiv
Top 0.1%
48.7%
Show abstract

Genetically informed conservation is critically important to ensure that interventions benefit the population of interest. In corals, preserving genetic diversity and accurate species identification are crucial for the sexual propagation. While various methods exist for species identification and measuring intraspecific variation, obtaining and analysing molecular data that enables rapid yet informed decisions on broodstock choice and progeny quality assurance remains challenging. Here we present a novel approach towards resource effective intraspecific genetic profiling by targeting extended genomic regions around ultra-conserved elements (UCEs). By sorting loci by parsimony informativeness and using a locus window size as small as 5000 bp upstream and downstream of the UCE, we identified a subset of 500 UCE-associated loci that can accurately resolve phylogenetic relationships among species and assess intraspecific variation with accuracy comparable to a whole-genome dataset, while. This method was validated using existing population genomic data from six species of staghorn coral (Acropora hyacinthus, Acropora tersa, Acropora pectinata, Acropora sp. "VI-3", Acropora kenti and Acropora cf. spathulata). The phylogeny produced by the UCE subset is congruent with the phylogeny based on complete data. With the moderate number and length of target genomic region sizes providing a balance between resolution and sequencing effort, this study provides a proof-of-concept approach towards developing fast, scalable, and cost-effective workflows using a real-time long-read sequencers such as Oxford Nanopore Technologies. The methodology has the broad potential to be applied to support genetic assessment across taxa where taxonomic uncertainty is common, improving confidence in experimental frameworks and conservation decisions.

4
Let the prey speak: Using PNA clamps to silence predator DNA in marine faecal diet studies

Polanowski, A. M.; Suter, L.; Deagle, B. E.; McInnes, J. C.

2026-07-08 molecular biology 10.64898/2026.06.22.733645 medRxiv
Top 0.1%
46.5%
Show abstract

DNA metabarcoding of faeces is a powerful, non-invasive method for assessing predator diets. However, when studying the diet of generalist predators, broad PCR primers are used to amplify the wide range of potential prey species and metabarcoding outputs are often dominated by sequences from the predator. While blocking primers can be used to reduce PCR amplification of predator DNA, they frequently cause partial predator suppression and unintended prey blocking. Peptide nucleic acid (PNA) clamps, offer a promising, underutilised alternative by binding strongly and selectively to predator DNA to block its PCR amplification. In this study we designed and validated a novel PNA clamp targeting the 18S rRNA gene to suppress bird and mammal predator DNA in dietary samples. We tested this clamp on tissue mixtures and faecal samples from three seabird and two seal species across temperate, subantarctic, and Antarctic regions. The PNA clamp substantially increased the proportion of prey reads recovered while maintaining consistent prey community composition across all predator species. Our results demonstrate not only the general effectiveness of PNA clamps over standard blocking primers, but also provide a powerful, broadly applicable new tool to improve the accuracy in DNA diet metabarcoding studies.

5
Low-Coverage Genome Sequencing Outperforms Target Enrichment Phylogenomics

Branstetter, M. G.; Freitas, F. V.; Benavides Silva, L. R.; Bossert, S.; Danforth, B. N.; Murray, E. A.

2026-06-09 evolutionary biology 10.64898/2026.06.05.730492 medRxiv
Top 0.1%
45.5%
Show abstract

Genome-scale data have transformed phylogenetic inference, yet most studies continue to rely on reduced-representation approaches that target a subset of loci to reduce cost and increase taxon sampling. Although effective, these methods require specialized laboratory workflows, constrain long-term data reuse, and may perform poorly with degraded DNA. Low-coverage whole genome sequencing (lcWGS) offers a streamlined alternative: shallow to moderate sequencing of complete genomes followed by bioinformatic extraction of loci of interest. Despite its promise, lcWGS has not been rigorously benchmarked against targeted enrichment using historical museum specimens. Here, we directly compared lcWGS and ultraconserved element (UCE) target enrichment across taxonomically diverse bee specimens collected between 1934 and 2021. Both data types were generated from the same Illumina libraries, enabling a controlled, head-to-head evaluation. Using standard UCE analytical pipelines, we quantified locus recovery, gene-tree support, and phylogenetic performance across sequencing methods and specimen age classes. We further assessed recovery of additional marker classes, including mitogenomes, BUSCO loci, and UCEs from a newly-designed, expanded probe set. Across all age categories, lcWGS consistently outperformed target enrichment, recovering more UCE loci and substantially longer alignments, with the largest gains observed in highly degraded specimens. Gene trees derived from lcWGS exhibited higher mean bootstrap support and greater topological concordance, translating into improved species-tree inference. In addition, lcWGS enabled recovery of markedly more non-target loci, expanding analytical flexibility beyond the original marker set. These results demonstrate that lcWGS not only matches but frequently exceeds the performance of targeted enrichment in museum-based phylogenomics, while providing broader genomic utility. As sequencing costs continue to decline, lcWGS represents a robust and forward-looking strategy for phylogenetic research, particularly in taxa with modest genome sizes and challenging DNA quality.

6
Optimization of High Molecular Weight DNA Extractions from Dried, Museum-Grade Insects Enables Long-Read Sequencing, Phylogenetics, and Methylation Profiling

Hartley, G. A.; Green, R. J.; Pauloski, N.; Tillquist, N. M.; Johnston, P.; Ord, S.; O'Neill, R. J.

2026-07-21 genomics 10.64898/2026.07.16.739004 medRxiv
Top 0.1%
44.5%
Show abstract

Developing an effective DNA extraction method that meets requirements for long-read sequencing of poorly preserved samples, such as museum specimens or ancient material, offers new opportunities for genomic analysis of endangered or extinct species for which samples are rare. However, these samples often yield degraded and highly fragmented DNA, rendering long-read sequencing infeasible for many specimens residing in museum collections. Herein, we demonstrate a protocol for successfully extracting DNA of sufficient quality for sequencing on the Oxford Nanopore Technologies long-read sequencing PromethION platform from a desiccated, museum-grade blue carpenter bee specimen (Xylocopa caerulea). We find the protocol is reproducible across specimens and yields high levels of long, endogenous X. caerulea-derived DNA, highlighting the utility of our method for enabling genomic studies of historical collections. From a single flow cell, we assembled the full-length mitochondrial genome and used this assembly to perform a phylogenetic analysis, accurately placing our X. caerulea specimen among related Xylocopa species, thus demonstrating the phylogenetic utility of long-read museomics. Using these long-read data, we analyzed native CpG methylation, finding endogenous methylation signals that correlate with genic and exonic sequences. This method expands the feasibility of genomic and epigenomic analyses from challenging samples, enhancing our ability to investigate the genomes of endangered and extinct species through archival resources.

7
Beyond DNA barcodes: an open-source workflow for recovering and organizing barcoded vouchers for ecological and evolutionary research

Feng, V.; Lin, H.-M.; Srivathsan, A.; Wang, H.; Lee, L.; Pedales, R.; Oberschmidt, D.; Meier, R.

2026-08-07 molecular biology 10.64898/2026.08.06.743289 medRxiv
Top 0.1%
41.7%
Show abstract

1. Most species are neither discovered nor named, let alone included in analyses that require biological information such as trait measurements, images, ecological information and genome-scale data. Specimen-level DNA barcoding can help discover many of these species rapidly, but everything beyond discovery requires vouchers organized into putative species. Yet, existing barcoding workflows lack efficient techniques for voucher recovery, creating a post-barcoding bottleneck that limits the ability of converting barcoded specimens into biological knowledge. 2. Here we present a low-cost, open-source workflow consisting of two stages. The first safeguards barcoded specimens by separating them from DNA extracts and transferring them from microplates into ethanol-filled glass vials. The second converts the resulting voucher collection into a searchable physical resource by linking barcode-derived molecular Operational Taxonomic Unit (mOTU) assignments to vial positions and enabling specimens to be sorted into putative species either manually or automatically using a newly developed open-access robot (SORTER). 3. We evaluated the workflow using 2,024 insect specimens distributed across 21 96-well plates. For the first stage, DNA separation and specimen transfer required approximately 15 minutes per plate. For the second stage, MOTUmapper generated retrieval coordinates in a few seconds, after which the 2,024 vouchers belonging to the 452 putative species could be recovered manually in 5 days or with SORTER in 5 hours. Throughout both stages, specimen identities remained linked to barcode sequences, metadata and storage positions. 4. Vouchers are the Rosetta stones of biology because they connect different kinds of data to the same specimens. By safeguarding these vouchers and making them searchable, the workflow converts barcode projects from one-time molecular surveys into reusable resources for ecological and evolutionary research.

8
The Metabarcoding Analysis Pipeline (MAP): Simple, accurate, and flexible metabarcoding

Prosser, S. W.; Bard, N. W.; Thompson, K. A.; Floyd, R. A.; Padhye, S.; Ozsahin, E.; Jafarpour, S.; Hebert, P. D. N.

2026-07-23 bioinformatics 10.64898/2026.07.22.740107 medRxiv
Top 0.1%
39.9%
Show abstract

Current metabarcoding pipelines are inflexible with respect to study design and are poorly suited to long-read sequence data. To address these limitations, we developed MAP, the Metabarcoding Analysis Pipeline, which is a sequence-to-answer workflow supporting the analysis of amplicons from highly multiplexed and replicated study designs. Although MAP can analyze amplicons of any length from any genetic marker, it includes several features tailored to long-read COI metabarcoding. MAP installs from a Docker container and requires only sequence data, a parameters file, and a reference library. It produces intuitive reports, enabling users to evaluate their data immediately after analysis. We validate MAP by showing that it generates biodiversity estimates that correspond closely to a ground-truth dataset of single-specimen DNA barcode data and by demonstrating that it outperforms alternative platforms for COI metabarcoding. MAP is free, open-source, and available from: https://github.com/cbg-innov/MAP.

9
Extraction workflow determines marker-specific recovery andreproducibility in leaf-litter eDNA metabarcoding

Weber, S.; Banerjee, P.; Scali, E.; Farrow, A. A.; Boren, A. M.; Russelk, W. T.; Gillespie, R.; Graham, N. R.; Roderick, G.

2026-07-16 ecology 10.64898/2026.07.15.738510 medRxiv
Top 0.1%
33.2%
Show abstract

Forest-floor leaf litter is a dynamic and structurally complex ecological transition zone and thus a promising substrate for terrestrial eDNA metabarcoding. Yet, extraction workflows for this heterogeneous matrix remain poorly standardized, especially in tropical systems, making it largely impossible to compare ecological functions across space, time, and taxa. To guide workflow selection across a series of selection criteria, including biological target, research question and practical considerations, we compared DNA extraction workflows for leaf-litter eDNA collected from 42 biological samples across seven different forest sites on Oahu, Hawaii. We evaluated four DNA extraction workflows: (1) Two low-volume approaches, with DNA extracted directly from 200 mg of homogenized litter using (i) CTAB or (ii) DNeasy PowerSoil(R); and (2) two high-volume approaches using PBS wash-based from 10 g of litter followed by (i) Centrifugation or (ii) Filtration. Taxonomic recovery from each workflow was evaluated with two COI primer sets targeting arthropods (ANML and shorter NoPlant), and one ITS marker targeting fungi. Results show that eDNA workflows tested here recovered site-level differences among forest-floor communities, but biodiversity recovery depended strongly on extraction workflow and marker. For low volumes, PowerSoil recovered the highest fungal richness (with ITS marker), and produced the most reproducible PCR-replicate profiles across markers, and required the least hands-on time, while CTAB was less expensive but required handling hazardous chemicals. For high volumes workflow, Centrifugation recovered higher arthropod diversity with ANML primer. Differences in community composition were nonetheless recovered by each method. At the same time, sampling sites explained more ASV-level compositional variation than extraction workflow across markers, showing that all workflows retained site-level ecological signals. Together, these results support a workflow framework in which extraction choice depends on target organism group, DNA state, reproducibility needs, and practical constraints.

10
The Barcode Inference Pipeline (BIP): From Sequencer Output to DNA Barcodes

Prosser, S. W.; Thompson, K. A.; Bard, N. W.; Floyd, R. A.; Ozsahin, E.; Hebert, P. D. N.

2026-07-24 bioinformatics 10.64898/2026.07.21.739876 medRxiv
Top 0.1%
32.4%
Show abstract

DNA barcoding involves the recovery of a DNA sequence for a target gene region from its source specimen. This process gains complexity when multiple sequences are recovered from a specimen, as is often the case when data are generated by high-throughput sequencers. This diversity can reflect both methodological artifacts (e.g., chimeras, PCR errors, sequencing errors, tag jumps) and real template diversity in the DNA extract (e.g., contamination, endosymbionts, NUMTs, parasites). To support analysis of the sequence data from three million specimens annually, the Centre for Biodiversity Genomics (CBG) has developed BIP, the Barcode Inference Pipeline. Compatible with all sequencing platforms, BIP processes .fastq files and returns both target DNA barcodes and non-target sequences. To generate results, BIP implements quality and size filtration, demultiplexing, primer trimming, chimera scanning, sequence error correction, OTU delineation, and sequence identification. When analysis targets the cytochrome c oxidase 1 (COI) barcode region, BIP also assigns each OTU to a known BIN or identifies its nearest neighbour BIN. As final output, BIP returns summary files ready for upload to BOLD or for other downstream analyses. They include a taxonomic assignment for each OTU, generated by comparison with a DNA barcode reference library. We describe BIPs flexibility and structure, then demonstrate its functionality by analyzing COI sequence data from 100K specimens. Because of its capacity to disentangle target and non-target sequences, BIP outperforms an alternative software package, ONTbarcoder, in several important ways. To ease access, installation, and functionality, BIP is provided as a Docker container (github.com/cbg-innov/BIP).

11
Metabarcoding replicate detection frequency tracks ddPCR copy number for cod and herring eDNA in ancient marine sediments

Banos Lara, E.; Holman, L. E.; Knudsen, S. W.; Bohmann, K.

2026-07-08 genetics 10.64898/2026.07.03.736335 medRxiv
Top 0.1%
31.6%
Show abstract

1. Detecting environmental DNA (eDNA) from rare or low-abundance aquatic species remains a major challenge, particularly when it is highly degraded, present at low concentrations, and dominated by DNA from non-target taxa. These challenges are further amplified in sedimentary ancient DNA (sedaDNA) studies, where thousands of years can degrade eDNA further, making the detection and quantitative interpretation of weak biological signals difficult. 2. Metabarcoding is commonly used to produce high-throughput community-level data from eDNA but is inherently compositional and influenced by amplification biases. Nonetheless, metabarcoding read abundance or PCR replicate detection frequency are increasingly used as proxies for relative DNA concentration, but their quantitative interpretation has rarely been evaluated against independent measures of absolute DNA abundance. 3. We used droplet digital PCR (ddPCR) to quantify mitochondrial DNA from Atlantic cod (Gadus morhua) and Atlantic herring (Clupea harengus) in 136 ancient eDNA extracts from Icelandic marine sediment cores spanning the last three millennia. We compared ddPCR copy number estimates with metabarcoding (18S) derived relative abundance and detection frequency, and evaluated whether temporal DNA trends corresponded with proxy reconstructed sea surface temperature (SST) variability. 4. We found that ddPCR-measured fish sedaDNA abundance was positively correlated with the proportion of metabarcoding PCR replicates for both Atlantic cod and Atlantic herring. Moreover, temporal trends in Atlantic herring DNA abundance were consistent with proxy reconstructed SST variability, supporting the ecological relevance of the molecular signal. 5. Overall, our results show that ddPCR-derived DNA concentrations and metabarcoding PCR replicate detection frequency capture consistent patterns in low-abundance fish sedaDNA from marine sediments. The observed agreement between approaches supports the use of PCR replicate detection frequency as a semi-quantitative proxy for low-abundance sedaDNA.

12
Charting the insect biodiversity of Crete: insights from a pilot metabarcoding survey

Koutsovoulos, G. D.; Sorg, M.; Hörren, T.; Buchner, D.; Bourlat, S. J.; Langen, K.; Trichas, A.; Leese, F.; Stamatakis, A.

2026-06-08 ecology 10.64898/2026.06.05.730060 medRxiv
Top 0.1%
28.6%
Show abstract

Among eukaryotes, insects are by far the most diverse organisms on Earth, yet their global decline threatens ecosystem stability. Understanding local and regional biodiversity patterns is critical for conservation planning, ecosystem management, and predicting responses to environmental change, but traditional surveys for assessing insect diversity (e.g., manual collection, morphological identification, and counting) are highly labor-intensive, time-consuming, and often require rare or simply unavailable dedicated taxonomic expertise. DNA metabarcoding offers an efficient, high-resolution alternative to assess insect communities. Here, we report on the first insect metabarcoding survey on Crete that spans two years of sample collection between 2021 and 2023 from a small area in Southern Central Crete in the context of a citizen science project. A total of 29 samples yielded 10,865 Exact Sequence Variants (ESVs), 10,516 of which were assigned to insects, covering 988 species, 900 genera, and 227 families across 14 orders. A comparison with the existing observation records reveals 406 potential newly-observed species and an estimated 690 unclassified species, indicating substantial cryptic diversity. Our results demonstrate that even small-scale sampling can unravel substantial insect diversity and highlight critical gaps in barcode reference databases. Our study demonstrates how DNA metabarcoding can accelerate biodiversity discovery and monitoring in understudied regions.

13
TRIDENT (Taxonomic Resolution and IDentification using Environmental dNa Traces): An Optimized Algorithm for Vertebrate Taxonomic Assignments in eDNA Metabarcoding, Integrating Molecular, Taxonomic, and Ecological Criteria

Haderle, R.; Jung, G.; Riou, M.; Ung, V.; Jung, J.-L.

2026-07-09 molecular biology 10.64898/2026.06.29.735257 medRxiv
Top 0.1%
22.7%
Show abstract

Environmental DNA (eDNA) metabarcoding has become a powerful approach for large-scale biodiversity assessment, yet taxonomic assignment remains one of its most critical error-prone steps. Current bioinformatic pipelines rely on molecular similarity searches against reference databases, but assignment accuracy is constrained not only by short marker length and database incompleteness, but also by fundamental limitations, including recent species radiations, incomplete lineage sorting, introgression, NUMTs, and the imperfect correspondence between genetic variation and species boundaries. Here, we present TRIDENT (Taxonomic Resolution and IDentification using Environmental dNa Traces), an automated and simple protocol designed to improve taxonomic assignments in eDNA metabarcoding. Initially developed for marine vertebrates, TRIDENT may be used with any barcode and integrates three complementary sources of evidence: molecular similarity (NCBI/GenBank and BOLD), curated taxonomic information (WoRMS), and ecological plausibility derived from biogeographic occurrence data (GBIF). The workflow sequentially constructs candidate taxon lists based on sequence similarity, expands them through taxonomic hierarchies, and filters them using spatial occurrence constraints. It further identifies possible taxa lacking reference barcodes and evaluates their plausibility through CO1-based similarity if data exist in BOLD. TRIDENT has been implemented as a source-available Python tool and tested using empirical eDNA datasets from marine vertebrates as well as simulated communities. Results demonstrate that the tool produces taxonomic assignments consistent with expert manual curation while substantially reducing processing time and attention errors caused by manual processing of large datasets. By combining molecular, taxonomic, and ecological criteria within a single framework, TRIDENT improves transparency and reproducibility and provides a robust and flexible solution strengthening confidence in taxonomic identifications in eDNA-based biodiversity assessments.

14
Full-length COI barcodes improve eDNA metabarcoding data denoising relative to mini-barcodes

Eisele, M. H.; Varusk, S.; Sammet, K.; Hakimzadeh, A.; Metsoja, M.; Tedersoo, L.; Alwutayd, K. M.; Arribas, P.; Andujar, C.; Emerson, B. C.; Anslan, S.

2026-07-03 ecology 10.64898/2026.07.03.736260 medRxiv
Top 0.1%
21.9%
Show abstract

Animal COI (mitochondrial cytochrome oxidase I) metabarcoding of environmental DNA (eDNA) is increasingly used to assess biodiversity in complex substrates such as soil. However, due to read-length constraints of second-generation sequencing platforms, mini-barcodes have been used instead of the full barcode region. Long-read sequencing technologies now enable the recovery of full-length barcode sequences, and are more commonly applied for studying microbes, but their use for metabarcoding the full-length standard COI barcoding region in animals remains limited. In this study, we compared three COI amplicon sets -- 313 bp, 660 bp, and 1,256 bp -- amplified from soil eDNA samples and sequenced using Illumina and PacBio platforms to evaluate their overall concurrence, the effectiveness of identifying nuclear mitochondrial DNA segments (NUMTs) and chimeras, as well as their respective taxonomic resolution. The long-read datasets exhibited a higher identification rate of NUMTs and true chimeras, suggesting that longer sequences improve the detection of noise in COI metabarcoding data, thereby reducing the occurrence of spurious taxa. Taxonomy assignment confidence was similar between the 313 bp and 660 bp datasets, whereas extending the amplicon beyond the standard COI barcode region (1,256 bp) reduced confidence, likely because longer reads extend into regions poorly represented in barcode reference databases. Despite substantially lower sequencing depth in the 660 bp dataset, per-sample OTU richness did not differ significantly from that recovered with the Illumina 313 bp amplicon set. Similarly, the relationships between samples were strongly correlated across the detected OTU communities, indicating consistent ecological interpretations between short and long amplicons. We conclude that the standard ~658 bp COI barcode is an optimal marker for soil animal metabarcoding from eDNA, balancing target recovery, artifact detection, taxonomic assignment and ecological interpretability. As COI eDNA metabarcoding becomes increasingly used in biodiversity assessment and is increasingly adopted in large-scale monitoring initiatives, this study provides methodological guidance for improving the robustness of soil animal community biomonitoring.

15
A chromosome-level genome assembly of a "living fossil", the tadpole shrimp Lepidurus arcticus (Pallas, 1793)

Strand, M. A.; Torresen, O. K.; Haga, J. A. R.; Danneels, B.; Skage, M.; Ferrari, G.; Tooming-Klunderud, A.; Hessen, D. O.; Jakobsen, K. S.

2026-06-08 genomics 10.64898/2026.06.04.730062 medRxiv
Top 0.1%
21.8%
Show abstract

We present the first chromosome-level reference genome for Lepidurus arcticus (Pallas, 1793), a freshwater crustacean with circumpolar distribution. L. arcticus belongs to the small order of freshwater Notostracan crustaceans that are representatives of the ancient group Branchiopoda. This group has a remarkable morphological stability and is frequently labelled "living fossils". Its ancient origin, streamlined genome (estimated to 0.11 Gb) and reproductive flexibility makes this a very interesting candidate for genomic studies. The haplotype-resolved assemblies are composed of two pseudo-haplotypes spanning 81.2 megabases (Mb) and 81.8 Mb, respectively, and each scaffolded into 6 chromosomes. Both haplotypes (hap) show high completeness and identical BUSCO scores of 98.3 for hap1 and hap2. The scaffold N50 length is 13.4 Mb for hap1 and 13.9 Mb for hap2, and k-mer completeness estimated from PacBio HiFi reads was 95.79% and 96.18%, respectively. The haplotypes display very low estimated genome-wide heterozygosity of 0.133%. The assembly contains 10901 (hap1) and 10910 (hap2) protein-coding genes. Repetitive elements comprised approximately 24-25% of each haplotype, with long terminal repeat retrotransposons representing the most abundant transposable element class at approximately 8-9%. Comparison with the near chromosome-level genome of Lepidurus packardi revealed substantial intrachromosomal rearrangements, despite similar chromosome numbers and chromosome sizes. Differences in transposable element content between L. arcticus and L. packardi were primarily driven by retrotransposons, particularly LTR and LINE elements. This reference genome provides a valuable resource for future population genomic studies and for investigating evolutionary stasis at the genome level.

16
Workflow for multiplex microsatellite panel development and sample preparation for robust amplicon sequencing of low-template and degraded DNA: validation for non-invasive genotyping in three large carnivore species

De Barba, M.; Boyer, F.; Baur, M.; Konec, M.; Pazhenkova, E.; Remollino, N.; Stoffel, C.; Boljte, B.; Miquel, C.; Skrbinsek, T.; Taberlet, P.; Fumagalli, L.

2026-08-21 ecology 10.64898/2026.08.20.745956 medRxiv
Top 0.1%
19.3%
Show abstract

High-throughput amplicon sequencing has transformed microsatellite (STR) genotyping by overcoming many of the limitations of fragment-length analysis, enabling more accurate, cost-effective, and standardized genotyping. Yet, protocols specifically designed for high-throughput sequencing (HTS)-based STR genotyping from low-template and degraded DNA remain scarce, despite the prevalence of these challenging sample types in ecological and conservation contexts. We present a methodology for the de novo development of robust STR multiplex panels together with a laboratory protocol for efficient and reliable STR genotyping by sequencing with low quantity and quality DNA samples. The protocol comprises (i) an automated bioinformatic pipeline to design large sets of short tetranucleotide markers optimized for multiplex amplicon sequencing of degraded and low-template DNA; (ii) guidelines for efficient in vitro optimization of multiplex amplification using directly low quantity/quality template DNA; and (iii) a library preparation procedure that improves detection of low-level allele signal while enabling quality assessment of STR amplicon sequencing under limiting DNA conditions. We demonstrate the approach by developing and validating STR panels for non-invasive genotyping of three large carnivore species: a 44-plex for the grey wolf (Canis lupus), a 41-plex for the Eurasian lynx (Lynx lynx), and a 30-plex for the brown bear (Ursus arctos). Multiplex performance was high, with [≥]91% of samples successfully genotyped at [≥]50% of loci (allele size range 28-110 bp across panels) and correctly assigned to known individuals, negligible levels of noise in the controls, and high discriminatory power (PIDsibs [≤]2.4 x 1e-12), also owing to sequence variation among same-length alleles at 15-50% of loci. The approach is broadly applicable to animal and plant species, a wide range of sample types, and large-scale analysis such as genetic monitoring. Our study reinforces the value of STR amplicon sequencing for ecological and conservation applications while highlighting the importance of marker design and laboratory workflows tailored to HTS-based genotyping for accurate and efficient implementation.

17
Microhaplotypes Improve Kinship Estimation in Heterozygous, Mixed-Ploidy Populations of Actinidia

Millar, T. R.; Koot, E. M.; Heywood, A.; Grande, A.; Thomson, S. J.; McCallum, J. A.; Wilcox, P. L.; Black, M. A.

2026-08-09 genetics 10.64898/2026.08.04.742852 medRxiv
Top 0.1%
18.6%
Show abstract

Over the past decade there has been increasing interest in the use of microhaplotype markers in autopolyploid taxa. This has been driven by theoretical and observed improvements in signals of allelic dosage, linkage, and heritability. Yet, to date there has been little investigation into the suitability of microhaplotype markers for estimating kinship. Here, we develop the theory of kinship estimation from microhaplotypes, introduce the MCHap microhaplotype caller for autopolyploid populations, and apply these methods to a highly diverse germplasm population of mixed-ploidy Actinidia (kiwifruit and relatives). We find that microhaplotype-based kinship estimates are generally superior to equivalent single nucleotide variant based estimates. This is because microhaplotypes minimize the coalescent signal among alleles which may bias estimates within the context of a recent reference population. Hence, kinship estimates from microhaplotypes more accurately capture the recent demographic history of a population. These findings are supported by both coalescent simulations and the analysis of real data. Our findings are relevant to organisms of any ploidy, but most actionable in highly heterozygous taxa such as Actinidia.

18
PhaseWY: A pipeline for haplotype phasing, sex chromosome identification and extraction of sex-limited sequences

Ellerstrand, S. J.; Churcher, A. M. J.; Kutschera, V. E.; Hansson, B.

2026-06-22 bioinformatics 10.64898/2026.06.17.732863 medRxiv
Top 0.1%
18.5%
Show abstract

Sex chromosomes are central to many ecological and evolutionary processes. Evidence has accumulated that sex chromosome systems vary extensively in age, turnover and transitions, motivating renewed efforts to study the diversity of sex chromosome systems across the tree of life. However, successful genomic detection of sex chromosomes depends on several factors, including the size and divergence time, background genetic diversity, and the number of sequenced females and males. In addition, technical challenges associated with sequencing and analysing the sex-limited Y/W chromosome remain. Here, we present PhaseWY, an automated Snakemake pipeline that uses whole-genome sequencing data from multiple female and male individuals to identify sex-chromosomal regions and extract the corresponding Y/W sequences. PhaseWY (i) detects sex differences in alignment depth, (ii) applies read-based and statistical haplotype phasing, (iii) identifies sex-linked regions using haplotype clustering, and (iv) subsets autosomal, X/Z- and Y/W-linked variants for downstream analyses. We applied PhaseWY to simulated data to benchmark factors influencing sex-linkage detection and successful extraction of Y/W-linked variants. To demonstrate its practical utility, we further applied PhaseWY to the neo-sex chromosome system in Alauda larks (Alaudidae) and performed a range of downstream analyses demonstrating the scope of applications of the PhaseWY output. We conclude that PhaseWY provides an easy-to-use and reproducible tool for population-genomic analyses in non-model organisms, with particular importance for advancing our understanding of sex-chromosome evolution.

19
A simulation-based method for genotype-environment association analysis

Sakamoto, T.; Yeaman, S.

2026-08-27 genetics 10.64898/2026.08.23.746561 medRxiv
Top 0.1%
18.2%
Show abstract

Genotype-environment association (GEA) analyses are widely used to identify loci underlying local adaptation by examining correlations between allele frequencies and environmental variables across a species' range. A major challenge for this approach is distinguishing true adaptive signals from spurious associations arising from population structure. Several methods have been developed to account for population structure, but these methods can suffer from reduced statistical power or increased false positives under some conditions. To address this, we introduce a new GEA method, termed SimGEA. In essence, SimGEA infers a neutral evolutionary model that reproduces the population structure observed in empirical data and uses this model to simulate neutral alleles. By applying the same GEA statistic to both the empirical and simulated data, SimGEA evaluates the significance of observed associations against neutral expectations that account for population structure. We compared the performance of SimGEA with that of existing GEA methods, including LFMM2 and BayPass, using simulations of local adaptation in two-dimensional space. We found that SimGEA consistently controlled the false discovery rate without substantially sacrificing statistical power across the scenarios examined. These results suggest that calibrating statistics using neutral simulations provides a robust and flexible approach for accounting for population structure in GEA analyses.

20
Insect COI barcoding data as an untapped resource for surveying Wolbachia symbioses

Nowak, K. H.; Buczek, M.; Marszałek, M.; Prus-Frankowska, M.; Valdivia, C.; Deng, J.; Shropshire, J. D.; Łukasik, P.

2026-06-25 ecology 10.64898/2026.06.24.734267 medRxiv
Top 0.1%
16.7%
Show abstract

O_LIDNA barcoding of the mitochondrial cytochrome c oxidase I (COI) gene is widely used to characterise insect diversity and distributions; however, its potential to reveal information on species interactions, including host-symbiont associations, remains largely unexplored. Here, we assess whether COI amplicon data can be used to identify Wolbachia - one of the most widely distributed bacterial symbionts known to profoundly affect their hosts biology. C_LIO_LIWe demonstrate that several commonly used invertebrate COI primer sets perfectly match many reference Wolbachia genomes, leading to frequent co-amplification. C_LIO_LIBy screening 7,901 individual-insect COI amplicon libraries obtained with the popular BF3-BR2 primer set, we detected Wolbachia sequences in over 35% of samples, revealing that co-amplification is indeed widespread. After removing low-abundance reads, Wolbachia detection based on COI amplicons showed over 90% agreement with simultaneously generated 16S-V4 rRNA amplicon data from the same specimens. The degree of agreement, however, varied depending on the thresholds used, among datasets and insect clades. C_LIO_LIFurther, we show that Wolbachia abundance inferred from COI amplicons correlated with their abundance in metagenomic datasets for 152 specimens, supporting the quantitative relevance of the signal. C_LIO_LIFinally, we find that Wolbachia COI sequences provide greater phylogenetic resolution than 16S-V4 rRNA data (mean pairwise genetic distance of COI sequences - 9.6%, 16S-V4 rRNA - 2.8%), and the reconstructed Wolbachia COI-based genotype network largely agrees with genome-based phylogenies. C_LIO_LICollectively, our results demonstrate that off-target Wolbachia sequences recovered from standard insect COI barcoding data may reliably detect symbiont presence, provide phylogenetic insight, and guide sample selection for metagenomics. Given the rapid expansion of global insect barcoding initiatives, these findings highlight an opportunity for cost-effective monitoring of their most prevalent bacterial symbionts, offering new perspectives on how host-microbe interactions may shape insect communities. C_LI